laya: add a reproduction command for every Apple Silicon benchmark number - #67
Open
cacheline999 wants to merge 4 commits into
Open
cacheline999 wants to merge 4 commits into
cacheline999 wants to merge 4 commits into
Conversation
…0 defaults Formatting only (88 columns, as the rest of the repository); no behaviour change.
… M4 in the recipe - The contract test of the first request after ready asserts that it succeeds with a valid answer. Its latency depends on how long the GPU has been idle and on the Mac (an M5 answered a first request without warmup in 132-143 ms), so it is left to the benchmarks; that the warmup ran before ready is tested separately. - requirements-mps.txt pins pytest, httpx2 and ruff; the recipe's Test section runs the ruff checks. - The recipe lists the reviewer's M4 among the Macs the tests ran on. - Test docstring points at tests/laya.
…cipe The benchmark README now has a table from each recipe claim to the command behind it. New scripts cover what ThinkFlowLab#30 measured with one-off probes: - paired.py --gap SECONDS (each request after that much idle, on a fresh connection) and --only; - lengths.py: the first request of new input lengths and the footprint they add, for one or two flag sets; it skips the lengths the warmup ran and stops at the window; - late_load.py: a checkpoint loaded while serving, directly and through the frontend; - fallback.py: Laya's fallback to the CPU under a lowered MPS memory limit; - release.py: in one process prepared by the worker's build_app, what torch.mps.empty_cache() gives back after new lengths. It releases once before the walk and reads every footprint --settle seconds after a release, because the footprint shows a release up to ~2 s late and the starting footprint varies by a few hundred MB between runs; where a release lands does not; - a loop of fresh starts alternating the worker with plain laya-serve. Every measuring script refuses a run on battery or above --max-load (late_load, fallback and release unless --feasibility); env.py holds the shared check and workload reader. Recipe changes from the reruns: a late load with --compile took 19-22 s (the ~70 s in ThinkFlowLab#30 was not reproduced, so the recipe gives no number for a load past the frontend's 60 s); the CPU fallback request took 30-73 s; every length up to the window is 454 lengths, 5.2 GB with the options and 3.7 GB without; a release brings the footprint back to about 2.95 GB. The heartbeat numbers had no script and are gone. Comments that ruff format had pushed onto closing brackets are back above their statements (syntax trees unchanged). No change to the worker's runtime code.
cacheline999
force-pushed
the
laya-review-followups
branch
from
October 4, 2026 09:17
32ff384 to
744b3b6
Compare
7 tasks done
…ince The recipe gave 0.7-1.1 s for plain laya-serve's first request after ready, from three runs on 2026-09-28. Seven fresh starts since, with the same torch, laya and macOS, took 0.2-0.4 s, both right after another run and after 30 s of idle, so the GPU's state left by a previous run does not explain it. The recipe now gives both and says the cause is not known. The worker's own first request (67-81 ms) is unchanged. The README says every script that measures refuses a noisy machine; report.py does not measure.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Follow-ups to #30:
Every performance number in the Apple Silicon recipe now has a command that reproduces it. Laya on Apple Silicon: worker, benchmarks and recipe #30
shipped scripts for the main comparisons, but several numbers in the recipe came from one-off probes
that were not in the repository: requests after an idle gap, the cost of new input lengths and the
memory they add, a checkpoint loaded while serving, and the CPU fallback. The benchmark README now
has a table from each recipe claim to the command behind it, and the missing pieces are added:
paired.py --gap SECONDS(each request after that much idle, on a fresh connection) and--only;lengths.py: first request of new input lengths and the footprint they add, for one or two flag sets;late_load.py: a checkpoint loaded while serving, its first request directly and through the frontend;fallback.py: Laya's fallback to the CPU under a lowered MPS memory limit;release.py: the memorytorch.mps.empty_cache()gives back after many new lengths, and whetherthose lengths are cold again afterwards. It prepares the model with the worker's own startup
(
build_app: options and warmup), exits if the model is not on MPS, walks lengths withlengths.py's walk, and reads each footprint--settleseconds after a release;laya-serve.
late_load.py,fallback.pyandrelease.pyrefuse a measured run on battery or under load likethe other scripts (
--feasibilityruns anyway); the check and the workload reader are now sharedhelpers in
env.py.Two statements had no command behind them and were changed instead: the recipe no longer gives
numbers for keeping the GPU busy between requests (it only says the worker does not do this), and
the ~70 s late load with
--compilecould not be reproduced (19–22 s alone, 17 s with a secondworker holding memory on the same GPU), so the recipe gives no number for a load past the frontend's
60 s and only says what happens then. The recipe's count of every length up to the window is now the
454 that the full-window command reaches (it said 477, from an earlier probe with shorter inputs).
Each is an A/B where the claim is a comparison (with the options against without, or a worker against
plain laya-serve). A rerun gives other numbers on another Mac or under other load; the comparison is
what should carry over. Rerunning everything on the M1 Pro changed these recipe statements (see Test
Result): the late load with
--compile(19–22 s; no number past 60 s), the CPU fallback request(30–73 s), every length up to the window (454 lengths, 5.2 GB with the options, 3.7 GB without),
what a release gives back (back to about 2.95 GB), and plain laya-serve's first request after ready
(0.7–1.1 s then, 0.2–0.4 s in seven later fresh starts, cause not known).
Readiness contract test (flaky on an M4 in the last review): it now checks that the first request
after ready succeeds with a valid answer, without a latency bound. That latency depends on how long
the GPU has been idle and on the Mac (an M5 answered a first request without any warmup in 132–143 ms),
so it stays in the benchmarks. That the warmup ran before ready is still tested.
ruff format drift: the Laya files are formatted with ruff 0.16.10 defaults (88 columns, as the rest
of the repository; the first commit is formatting only, with every file's syntax tree unchanged), and
requirements-mps.txtpins ruff, pytest and httpx2. Comments that the formatter had pushed ontoclosing brackets are back above their statements, so the two
noqa: SIM115suppress again.Stale docstring in
tests/laya/test_worker.py; the reviewer's M4 is listed in the recipe.No change to the worker's runtime code. A CI job for
tests/layawas suggested in the review; happy toadd one in a separate PR if you want it.
Test Plan
System1-Omni Version / Commit:
62dc52aon top of58b8cberuff format --checkandruff check --select E4,E7,E9,Fon the Laya files, ruff 0.16.10.PYTHONPATH=src python -m pytest tests/laya, and withLAYA_CONTRACT=1; contract tests on MPSwithout and with
--compile --weights fp16.script's flags (a misspelled flag fails it).
mkdocs build --strict.Test Result
All reruns on the M1 Pro (16 GB). None was on an idle machine on mains power, so they were run as
feasibility; what each shows is that the command runs and whether the recipe's effect is there. Theconditions column says which session a row comes from.
release.pyempty_cache()after 100 new lengths, with the options--compileChecks: ruff format and check clean; 98 unit tests passed, 118 with
LAYA_CONTRACT=1; contract on MPS20/20 plain and with the options; strict docs build passes.
Self-review
Several self-review passes ran before this description; each issue below was reproduced (a test that
failed, or a rerun) before it was fixed.
lengths.pynow skips the lengths the warmup ran (read back from the worker) and stops at the window;the full-window walk covers 454 lengths, 57–512 tokens.
late_load.pywaits up to--timeout, follows up a load cut off by a 504 or a timeout only if thecheckpoint became resident, refuses a
--latethe worker already serves, and documents every field.release.pyprepares the model the way the worker does, exits off MPS, keeps the traceback of otherfailures, reads W1 by id, and walks lengths with
lengths.py.Every script that measures refuses a noisy machine;
fallback.pysays why its memory probe failed.Two measurements did not hold and are not in the recipe: a release giving back "about half" and then
"about 60%" came from a starting footprint that varies by a few hundred MB between processes and from
reading the footprint before the release had landed (
release.pynow reads--settleseconds later);a
torch.mps.synchronize()before the release made no difference in an A/B (it waits 0.004 ms aftera request). The ~70 s late load could not be reproduced, with or without a second worker.
I have reviewed the full diff and addressed the issues I found.
I have checked that the change follows the project's architecture and stays focused on the stated purpose.
I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.